G3: Genes|Genomes|Genetics
Preprints posted in the last 7 days, ranked by how well they match G3: Genes|Genomes|Genetics's content profile, based on 35 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.
Show abstract
Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.
Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.
Show abstract
Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.
Hussain, T.; Anothai, J.; Nualsri, C.; Ali, A.; Khomphet, T.
Show abstract
Drought stress is the major yield limiting factor in upland rice production where the moisture availability is highly variable. Understanding and evaluating how upland rice responds to drought stress is critical to improving resilience and yield stability. In this study performance of sixteen upland rice varieties were evaluated under non-stressed, moderately stressed and highly stressed conditions. Drought stress was introduced by irrigating upland rice at 70% and 50% field capacity (FC) whereas non-stress treatment was irrigated at 100% FC. Irrigation in moderately stressed and highly stressed conditions was also withheld for six days at lateral crop stages to observe temporary wilting by inducing a stress interval. Data on agronomic traits of upland rice was collected in three experimental replications. Results indicated that performance of upland rice varieties was significantly altered under stress conditions and highest performance was observed under non-stressed conditions. Yield losses for short duration and long duration varieties ranged 35-60% and 24-62% under moderate stress whereas it ranged 43-78% and 52-73% under highly stressed conditions, respectively. Overall varieties Dawk Kha, Khao/ Sai and Dawk Pa-yawm, indicated higher stability under stressed conditions therefore, these long duration varieties could be used for obtaining better yields under diverse agroclimatic conditions and under unpredicted weather patterns. Short duration Ma-led-nai-fai and long duration Goo Meung Lung and Bow Leb Nahag could be used for acquiring traits for higher tillering and panicle bearing capacity. Short heighted varieties such as Jao Daeng, Sahm Deuan and Ma-led-nai-fai could be used in breeding for short heighted new varieties to overcome lodging concerns. Strong significant association of GMP, STI, MPRO, MHAR with grain yield under non-stressed, moderately stressed and highly stressed conditions indicated that these indices were appropriate for their use as selection criteria for drought resilience.
Saqib, M.; Chen, F.; Mistri, D. K.; Tan, L.; Wright, N.; Sarver, D. C.; Anders, R.; Aja, S.; Wong, G. W.
Show abstract
Trisomy 21 or Down syndrome (DS) affects multi-organ systems across the lifespan. The presence of an extra chromosome, along with genome dosage imbalance due to triplicated genes, contributes to the DS phenotypes. Of the DS mouse models, few are aneuploid with a freely segregating extra chromosome. We previously showed that the aneuploid Ts65Dn mice exhibit metabolic deficits consistent with the metabolic profile of DS. However, the genotype-phenotype relationships in Ts65Dn mice are complicated by the presence of triplicated genes unrelated to human chromosome 21 (Hsa21). To address this issue, we leveraged a refined model, Ts66Yah, where the extra triplicated genes in Ts65Dn have been removed. Deep phenotyping and multi-omics analyses showed that Ts66Yah mice develop pronounced and widespread metabolic disturbances. Despite sexual dimorphism in weight gain, body temperature, lipid and lipoprotein profiles, hepatic injury and adipose fibrosis, both male and female Ts66Yah mice share a common phenotype of pronounced glucose intolerance and insulin resistance, reduced mitochondrial respiratory capacity in visceral fat, altered serum inflammatory cytokine profile, and dysregulated serum and liver metabolomes. Pan-tissue transcriptomes also reveal signatures of immune activation, disrupted metabolic processes and cellular respiration, altered cytokine signaling, enhanced oxidative stress, and extracellular matrix remodeling. These combined changes across tissues disrupt metabolic homeostasis more severely in Ts66Yah than in Ts65Dn mice. Several phenotypes, including glucose intolerance, insulin resistance, tissue fibrosis, and oxidative stress were further exacerbated by an obesogenic diet. This foundational data establishes Ts66Yah as a valuable reference model for the mechanistic and comparative study of metabolic dysfunction in DS.
Duffin, P. J.; Ruggeri, M.; Conn, T.; Baums, I. B.; Blanco-Pimentel, M.; Bosch, P.; Carne, L.; Danser, N.; Montoya-Maya, P.; Morikawa, M.; Muller, E. M.; Winters, R. S.; Baker, A. C.; Cunning, R.; Dahlgren, C.; Parkinson, J. E.; Kenkel, C. D.
Show abstract
Genomic signatures can provide key insight into the evolutionary history and remaining adaptive potential of threatened populations. As demographic decline erodes both diversity and the processes maintaining it, understanding how remaining variation is distributed becomes increasingly important for conserving species like the staghorn coral, Acropora cervicornis, a foundational but critically endangered Caribbean reef-builder. We analyzed 46 high-coverage A. cervicornis genomes from 10 locations across the tropical western Atlantic to evaluate neutral and adaptive structure, genomic diversity, demographic history, inbreeding, and connectivity, and generated a regional haplotype reference panel for future genomic monitoring. Genome-wide analyses recovered recurring regional substructure, but differentiation was modest and partly explained by isolation-by-distance and spatial variation in effective migration. Subpopulations had similar levels of genomic diversity, shared demographic history, and limited evidence of local adaptation. These patterns support interpreting sampled Caribbean populations as a single evolutionarily significant unit (ESU) containing multiple regional management units (MUs), rather than as deeply divergent evolutionary lineages. Despite substantial retained variation and low current inbreeding, estimated contemporary effective population size was small, suggesting an increased vulnerability to the effects of drift as demographic collapse continues, especially if structure is reinforced by isolated management. Together, our findings emphasize the urgent need for interventions that preserve and enhance genomic diversity, including risk-managed assisted gene flow. Supported by the haplotype reference panel developed here, these strategies will require coordinated efforts across regional entities to conserve and restore A. cervicornis as a jointly managed, single ESU.
Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.
Show abstract
Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.
Mansoor, R.; Minhas, A. S.; Thomas, A.; Mansoor, A. A.; McCambridge, A. H.; Dilts, C.; Eshak, J.; Govani, D.; Nylin, B.; Trinidad, J. C.; Kanaan, A. Y.; Kara, E.; Fielder, A.; Fielder, I.; Iglendza, A.; Mukatash, Y.; Pumnea, B.; Menzel, M. M.; Shabazz-Henry, A. L.; Niepielko, M. G.; Gao, M.
Show abstract
The QxxR motif is evolutionarily conserved within DEAD-box RNA helicases, including Drosophila Me31B and human DDX6, which post-transcriptionally regulate gene expression during animal development. A pathogenic H372R substitution (QxHR to QxRR) in the QxxR motif of human DDX6 has been associated with various developmental defects, but how this motif contributes to DDX6-family protein function remains unclear. Here, we used Drosophila Me31B as an in vivo model to investigate the QxxR motifs developmental role. We generated a Drosophila strain carrying the corresponding H333R missense mutation in Me31B and characterized its effects on female fertility, embryonic viability, germline development, and Me31B-associated molecular pathways. The me31BH333R mutation reduced female fertility in a gene dose-dependent manner, with homozygous mutant females being sterile. Embryos from the mutant females also exhibited primordial germ cell defects. Despite these developmental phenotypes, the me31BH333R mutation did not significantly alter Me31B protein abundance, global ovarian transcriptome or proteome profiles, or representative germ plasm mRNA and protein localization. In contrast, bait-normalized IP-MS analysis revealed altered enrichment of selected Me31B-associated proteins, including increased association of known Me31B interactors Trailer hitch (Tral) and Ypsilon Schachtel (Yps). These findings establish Me31BH333R as an in vivo model for investigating the conserved QxxR motif and suggest that disruption of this motif compromises development not through broad changes in gene expression, but potentially through altered composition or regulation of Me31B-containing ribonucleoprotein complexes.
Kwon, H. R.; Rackley, A.; Olson, L. E.
Show abstract
Autosomal dominant gain-of-function mutations in platelet-derived growth factor receptor beta (PDGFRb) cause overgrowth of the skeleton and other connective tissue in Kosaki overgrowth syndrome. However, the target cell type and signaling pathways underlying PDGFRb-driven overgrowth are unknown. Normal postnatal growth is controlled by pituitary-secreted growth hormone (GH), which activates the STAT5 transcriptional factor to upregulate insulin-like growth factor 1 (IGF1). To investigate the role of the GH-STAT5-IGF1 pathway in PDGFRb-related overgrowth, we generated mice with a PDGFRb gain-of-function mutation in skeletal and fibroblast lineages, which resulted in STAT5 activation and gigantism. Conditional deletion of Stat5ab in connective tissue lineages rescued skeletal overgrowth and keloid-like fibrosis in the skin. Conditional deletion of GH receptor (Ghr) did not rescue overgrowth, indicating the physiological activator of STAT5 is not required for overgrowth. However, deletion of Igf1, the STAT5 target gene, and its receptor, Igf1r, in connective tissue, rescued the overgrowth phenotype. These findings demonstrate a GHR-independent STAT5-IGF1 signaling pathway in mutant connective tissue cells, which mediates PDGFRb-driven overgrowth in mice and potentially in humans with similar PDGFRB mutations.
Werner, A. P.; Sachithanandham, J.; Akin, E.; Talukdar, S.; Pinsley, M.; Pekosz, A.
Show abstract
H5N1 clade 2.3.4.4b avian influenza A viruses pose a significant threat to wild animal populations, domesticated animals, and potentially, the human population. For H5N1s to infect and transmit among mammalian species, mutations for improved utilization of mammalian receptors and enhanced replication at the lower temperatures of the upper respiratory tract need to be acquired. A human H1N1pdm09-like virus was compared to H5N1 genotypes B3.13 and D1.1 for replication at 33{o}C, 37{o}C, and 39{o}C - temperatures consistent with the upper and lower respiratory tract in humans, and dairy cow udder tissue. All H5N1 viruses had increased plaque sizes on MDCK cells at 37{o}C and 39{o}C compared to H1N1pdm09. In primary, differentiated human nasal and bronchial epithelial cultures, all H5N1 viruses show restricted infectious virus production compared to H1N1 at 33{o}C. While H5N1 D1.1 also showed restricted replication at 37{o}C and 39{o}C, the H5N1 B3.13 replicated to nearly equivalent titers as H1N1pdm09. All H5N1 viruses demonstrated similar cell tropism in cells from the upper and lower respiratory tract, infecting more ciliated than non-ciliated cells relative to H1N1pdm09. H1N1, H5N1 B3.13 D1.1 infection induced similar innate immune factors, with nasal epithelial cells producing higher levels compared to bronchial epithelial cells. These data suggest that genotype B3.13 and D1.1 H5N1 viruses show different temperature dependent replication patterns compared to H1N1pdm09.
Santoyo, G.; Flores, A.; Castelan-Sanchez, H. G.; Valenzuela-Ruiz, V.; de los Santos-Villalobos, S.; Mitra, D.; Babalola, O. O.; Schoebitz, M.; Orozco-Mosqueda, M. d. C.
Show abstract
Plant growth-promoting bacterial endophytes represent a sustainable strategy for enhancing agricultural productivity while reducing reliance on synthetic fertilizers and pesticides. This study focused on the genomic and functional characterization of two endophytic bacterial strains, R11F and R19M, isolated from bean and maize roots, respectively. Comparative analyses based on 16S rRNA gene sequences, average nucleotide identity (ANI), and genome-to-genome distance calculations (GGDC) classified both isolates as Pseudomonas palleroniana. Comparative genomic analyses revealed highly conserved genomes containing genes associated with plant colonization, phosphate solubilization, stress adaptation, heavy metal resistance, and hydrocarbon degradation. Genome mining further identified 17 and 18 biosynthetic gene clusters (BGCs) in R11F and R19M, respectively, including non-ribosomal peptide synthetases (NRPS), pyoverdine, NRP-metallophores, RiPP-like compounds, arylpolyenes, {beta}-lactones, terpenes, NAGGN, and hydrogen cyanide. Strain-specific BGCs associated with syringomycin and viscosin biosynthesis were identified in R11F, whereas R19M harbored clusters related to asplenin and kolossin biosynthesis. In vitro assays confirmed indole production, phosphate solubilization, and siderophore production, as well as the ability of both strains to grow in nitrogen-free medium. Both strains significantly inhibited the growth of Fusarium oxysporum, Phytophthora cinnamomi, and Colletotrichum gloeosporioides. Furthermore, plant inoculation assays demonstrated host-dependent growth promotion, with R11F showing the most consistent improvements in plant growth parameters in tomato, wheat, and lentil. Overall, the integration of comparative genomics and experimental validation demonstrates that P. palleroniana R11F and R19M possess complementary traits associated with plant growth promotion, pathogen suppression, saline stress adaptation, and bioremediation.
Pollenz, R. S.; Davenport, M.; Ruiz-Houston, K. M.
Show abstract
Phage D29 infects Mycobacterium smegmatis mc2 155 and has a non-canonical lysis cassette that encodes two endolysin proteins (Lysin A and Lysin B) and a single two transmembrane domain (TMD) protein, LysA2a similar to F1 cluster phage LysF1a. A 1TMD LysF1b homolog, LysA2b, is encoded by a gene found downstream of the tape measure. Exogenous expression of both LysA2 proteins in tandem is a cytotoxic to M. smegmatis. Deletion of lysA2a produces phages that are lysis competent with a 10-minute triggering delay and 30% plaque size reduction. Deletion of lysA2b results in severe lysis defects manifest by 70% reduced plaque size, delayed lysis timing and reduced burst size. Deletion of both lysA2 genes results in phages that are viable and show lysis phenotypes like the lysF1b deletion. Genetic complementation of lysA2b deleted phage with the lysF1b gene fully complements the lysis phenotypes but alters the triggering time to that of an F1 cluster phage. Energy poisons trigger lysis prematurely in all phages with lysA2 gene deletions. Lysis recovery mutants (LRM) isolated from phages lacking the lysA2b genes generate wild type plaque size and have point mutations that map to TMD1 or the C-terminal region of the lysA2a gene. LRMs isolated from phages lacking both lysA2 genes show premature lysis and have mutations that all map to residue C31 of a novel lipoprotein (gene 64). Deletion of gene 64 does not change wild type D29 lysis phenotypes or rescue the lysis defects of any of the lysA2 mutants. A fitness/competition assay shows that loss of the lysA2 genes imposes a substantial competitive fitness cost. These finding support a lysis regulatory network model where the 2TMD protein is maintained in an inactive state until activated by its cognate 1TMD lysis regulator and the lipoprotein has accessory function that may enhance lysis efficiency.
Alquicira-Hernandez, J.; Dorans, E.; Tomofuji, Y.; Nathan, A.; Raychaudhuri, S.
Show abstract
Single-cell technologies enable linking disease-risk variants to gene regulatory effects in specific cell-state contexts. However, most so called "single-cell eQTL" studies use a "pseudobulking" strategy to identify expression Quantitative Trait Loci (eQTLs), obscuring subtle dynamic regulatory effects of disease alleles. Here, we propose Dynema (Dynamic eQTL mapping in single cells) for fast and accurate genome-wide mapping of context-dependent and independent eQTL effects at true single-cell resolution. To identify eQTLs, Dynema uses a Poisson model with cluster robust variance estimators (CRVEs) to account for correlation of single-cell profiles from the same individual. In contrast to other common methods, Dynema achieves statistical calibration and scales to genome-wide analysis in large single-cell datasets in realistic timeframes. We applied Dynema to two independent T cell datasets and identified reproducible cell-state-dependent eQTL effects. Some cell-state-dependent eQTLs are missed by pseudobulking approaches, and many others are conditionally independent from lead eQTL effects. We show that TSPAN32 and other autoimmune loci colocalize with cell-state-dependent eQTLs. Mapping context-dependent eQTLs at single-cell resolution enables the definition of the molecular effects of complex disease alleles.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.
Show abstract
Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.
Bowness, J. S.; Bernal Martinez, A.; Barinka, J.; Schulte-Schrepping, J.; Renders, S.; Waclawiczek, A.; Leppa, A.-M.; Trumpp, A.; Raffel, S.; Haas, S.; Velten, L.
Show abstract
To sustain blood formation, hematopoietic stem and progenitor cells (HSPCs) coordinate a multitude of cell biological processes, from cell cycle control and stress responses to lineage priming. While many genetic regulators of high-level HSPC function have been identified, how HSPCs coordinate more basal cell biological programs, and how such programs relate to stem cell function, remains incompletely understood. Here we use Perturb-seq to profile the transcriptional consequences of targeting 520 genes by CRISPRi in primary mouse HSPC cultures. We developed an analytical strategy to separate perturbation-induced changes in cell-state abundance and clonal heterogeneity from cell-state-local transcriptional effects. From these local perturbation signatures, we identified 19 gene regulatory programs (GRPs) that are defined by co-regulation in response to genetic perturbation, in contrast to co-expression or human curation, and align well with cell biological processes. By decomposing gene expression data from functional and clinical studies into program activity, we show that GRP activities associate with, and predict, phenotypes such as clonal output after transplantation, as well as survival and drug response in retrospective acute myeloid leukemia (AML) cohorts. Together, our study establishes perturbation-derived co-regulation programs as an interpretable framework for linking genetic regulators, cell-biological processes and stem-cell-associated phenotypes.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.
Show abstract
Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.
Mannava, S.; Ramkumar, V.; Murthy, G.
Show abstract
Introduction Hearing loss (HL) affects over 1{middle dot}5 billion people globally and India shares a disproportionately high burden including Disabling Hearing Loss (DHL). HL affects an Individual socio-economically, but there are limited studies on the broader societal economic consequences of HL in India.Methods Using Cost-of-Illness (COI) approach, we studied the societal economic burden of HL in India. This study uses epidemiological and macroeconomic data and modelling to estimate the loss of Gross National Income (GNI) due to HL and DHL across three economic pathways. Uncertainty is evaluated using deterministic and Probabilistic Sensitivity Analyses (PSA).Results The model estimates that there are in India, 289 million and 85{middle dot}9 million people with HL and DHL respectively. Direct Loss of GNI and Indirect Loss of GNI (Caregiver burden) are estimated as INR 4,648{middle dot}4 billion (USD 55{middle dot}6 billion) and INR 3,268 billion (USD 39 billion) respectively. The Loss of GNI due to Low Education amongst those with HL is estimated as INR 1,041{middle dot}9 billion (USD 12{middle dot}45 billion).Discussion Economic burden of HL is presented across three pathways with Direct Loss of GNI due to DHL being the greatest. It also presents age stratified caregiver economic burden. The findings of the study help in estimating similar cost pathways, advocacy, and policy decisions towards reducing HL prevalence in India and LMICs. This study also highlights the need for India specific estimations related to the HL attributable low education, state-wise disaggregates, and prevalence studies. Funding This study has not received any funding.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.